Tag
15 articles
GPT-6 Astra demonstrates a major leap in spatial reasoning on robotics benchmarks, outperforming competitors in dual-arm robot tasks.
Google Research introduces ToolGrad, an innovative framework that generates high-quality tool-use datasets with a 99.8% pass rate on ToolBench, outperforming traditional methods.
OpenAI's GPT-6 Astra receives mixed reviews on benchmarks, but its superior efficiency on ARC-AGI-3 has prompted AI expert François Chollet to accelerate his AGI forecast.
Google DeepMind is pioneering a double-blind evaluation process for AI models using Confidential Space to ensure unbiased benchmarking. The initiative, tested with a Gemini Flash Lite model and the Singapore AI Safety Institute, could redefine how AI systems are assessed.
Deepseek's new experimental multimodal model, V4-Flash-Vision-Exp, rivals Opus 4.8 on agent benchmarks by combining text and image understanding capabilities.
Alibaba's Qwen3.8 Max surpasses Claude Opus 4.8 with a 10-point jump, but Kimi K3 still leads with 25% better performance at lower cost.
Anthropic's Claude Opus 5 outperforms competitors like Claude Fable 5 and GPT-5.6 Sol in key benchmarks while offering significantly lower costs.
OpenAI's GPT-5.6 Sol nearly matches Claude Fable 5 on benchmarks while costing one-third as much, signaling a major shift in AI pricing and performance dynamics.
Anthropic's Claude Fable 5 leads industry benchmarks in finance, law, and medicine, but its $3.48 per task price tag is over 100 times more than competitors like DeepSeek V4 Pro.
OpenAI introduces Genebench-Pro, a new benchmark for evaluating large language models' biological and medical understanding capabilities. The tool aims to advance AI applications in healthcare and scientific research.
Anthropic's Fable 5 briefly outperformed OpenAI's GPT-5.5 before being shut down by the U.S. government, sparking speculation about national security concerns and AI regulation.
GPT-5.5 tops AI benchmarks but still hallucinates frequently, and its API cost has risen by 20%.